Every outage, caught in seconds.
Carbon checks your endpoints from twelve regions, pages the right engineer, and writes the status update, so the incident is over before the first support ticket lands.
Checkout API returning 502s in us-east-1
checkout-api us-east-1
-
Detected 03:14:02
3 of 12 regions returned 502 twice in a row
-
Paged 03:14:50
Priya acknowledged from Slack in 48 s
-
Root cause 03:17:31
inventory-svc timing out after a bad deploy
-
Resolved 03:20:14
Rollback finished, status page updated
Trusted by teams who get paged
One bill instead of four.
Uptime, incidents, logs, and status pages are one product with one price. Ingest everything, sample nothing, and stop paying per seat for the people who only ever read the status page.
- Ingest up to
- 80× more logs
- for the same budget
- or cut the bill by
- 92%
- at your current volume
- 1 TB
- logs per month
- 200
- monitors on 30 s checks
- 5
- engineers on call
Your current vendors
approx. $4,860 per month
Carbon
$389 per month
An estimate. Assumes annual billing, 1 TB of logs a month with 30-day retention, 200 monitors on 30-second checks, one status page, and a five-person on-call rotation. Your numbers will differ; the ratio rarely does.
Uptime monitoring
Explore uptime monitoring-
- Response
- Screenshot
- Timeline
Captured 03:14:02 · Frankfurt, DE · 1440×900
A screenshot of every failure
When a check fails, Carbon records the response and photographs the page, so you see what your customer saw.
-
- Traceroute
- MTR
- SSL
- cURL
Frankfurt, DETraceroute and MTR on every timeout
A timeout is not an answer. Carbon runs traceroute and MTR from the probe that failed, so a flapping transit hop looks like one.
-
api.example.com · 90 days 99.98% uptime90 days agoToday
Ninety days of honest uptime
Every check from every region, kept for ninety days and drawn as bars your customers can read on the status page.
Incident management
Explore incident management-
Carbon APP 03:14
Incident opened from monitor "Checkout API · POST /orders"
Checkout API returning 502s in us-east-1
Error rate 14.2% over the last 60 s from 3 of 12 regions. Paging Priya Natarajan (primary).
Acknowledge Resolve Snooze 30 minPaged where you already are
The page lands in Slack, Teams, SMS, or a phone call, with acknowledge and resolve on the message itself.
-
On-call this week Platform · follow the sun
-
Priya Natarajan
Primary · until Thu 09:00
-
Tomás Ferreira
Secondary · escalates after 5 min
-
Engineering manager
Escalates after 15 min
Schedules that follow the sun
Rotations, overrides, and a three-step escalation ladder. Nobody gets paged twice for one outage.
-
-
Incidents Last 24 hours
Similar incidents merge
Ten alerts fire at once when a database goes down. Carbon opens one incident and keeps your phone from ringing ten times.
Log management
Explore log management-
Query Sampling off2,567,345 rows in 0.7 s SQL · PromQL · Drag and drop
Query every line, sample nothing
SQL, PromQL, or drag and drop over raw logs at any volume. The answer comes back in under a second because nothing was thrown away.
-
Live tail · checkout-api live
Add drop rule
Stop ingesting lines like this one. Nothing matching it is billed.
Drop rules at the edge
Right-click a noisy line and add a drop rule. It stops being ingested, and it stops being billed.
-
HTTP 5xx rate · checkout-api Anomaly detected09:0010:0011:0012:00
Anomalies, not thresholds
Carbon learns the shape of each metric and alerts on the spike, so you never tune a threshold at 3 am.
Status pages
Explore status pages-
Northwind Status
All systems operational
Your brand, your subdomain
A status page on status.yourdomain.com, styled with your colours and logo, and fully customisable with CSS.
-
Get status updates
We email you whenever Northwind opens, updates, or resolves an incident.
Customers subscribe to the parts they use
Email and RSS subscriptions per component. An incident on webhooks never emails the people who only use the dashboard.
-
p95 response time · api p95 184 msMonTueWedThuFri
Response time, in public
Publish p95 response time next to uptime. It is the number your customers ask about second.
Pages land where your team already is.
Alerts go to the chat tool, the incident tool, and the phone you already carry. Acknowledge from any of them and Carbon stops paging everyone else.
Carbon APP 03:14
Incident opened from monitor "Checkout API · POST /orders"
Checkout API returning 502s in us-east-1
Error rate 14.2% over the last 60 s from 3 of 12 regions. Paging Priya Natarajan (primary).
Don’t take our word for it.
Teams from two-person startups to public companies run their on-call on Carbon. This is what they say when nobody from sales is in the room.
-
Went from zero to logs, uptime, and a status page in an afternoon. The bill is a fifth of what we paid for two of those things last year.
Conor Walsh
@cnrwalsh
-
Our domain expired at 2am. Carbon paged me about the certificate six days earlier and I ignored it. That one is on me.
Quinn Ferrara
@qferrara
-
Switched from a status page vendor over a weekend. Custom domain on the free plan, which nobody else does. Looks better than ours did.
Tian Zhou
@tianzhou
-
The traceroute-on-timeout thing has ended three arguments with our CDN this quarter alone.
Darren Pinder
@dpinder
-
I monitor one Ubuntu box for a side project. Log alerts, downtime, Slack pings, S3 archive. Still free. I keep waiting for the catch.
Kostya Melnyk
@kmelnyk
-
One tool, one bill, one place to look when it breaks. Our on-call rotation stopped complaining about the on-call rotation.
Marisol Reyes
@marisolr
-
Incident merging is the feature nobody demos and everybody needs. A database blip used to be twelve pages. Now it is one.
Jules Okafor
@julesok
-
Asked support a question at 23:40 on a Sunday. Got an answer from an engineer at 23:52. That is the whole review.
Sasha Lindqvist
@sashalq
-
We ingest everything now. No sampling, no 'top 1000 rows'. The query that used to time out takes 700 ms.
Priya Natarajan
@priyan
-
The screenshot of the failed check is the first thing I paste into the incident channel. Ends the 'works for me' phase instantly.
Rowan Blake
@rowanblake
-
Moved forty monitors over with the Terraform provider before lunch. The import wrote the resources itself. I did not expect that to work first time.
Eli Marchetti
@elimarch
-
Our status page finally matches our brand and the uptime bars are the real numbers, not a marketing graphic. Customers noticed.
Noor Haddad
@noorhdd
Sleep through the next one.
Set up your first monitor in two minutes. The free plan pages you, and stays free.
Start monitoring for free or book a demo
- Virginia 21 ms
- Oregon 58 ms
- São Paulo 112 ms
- Dublin 74 ms
- Frankfurt 82 ms
- Stockholm 91 ms
- Mumbai 188 ms
- Singapore 204 ms
- Tokyo 156 ms
- Sydney 221 ms
- Cape Town 240 ms
- Toronto 29 ms